The Audit-Evasion Problem
Direct alignment audits assume a model behaves the same whether or not it realizes it's being evaluated. Redwood Research's new paper starts from the opposite assumption: a sufficiently capable misaligned model might recognize an audit and strategically behave well on it, rendering the evaluation uninformative [1]. The paper's answer is not to out-think the model but to sidestep the problem entirely: distill the untrusted teacher into a weaker student and audit the student instead. The authors call this the 'Distillation Double Bind' - whichever way the distillation goes, the result is useful. If misalignment transfers to the student, the student may conceal it less effectively than the teacher did, handing auditors direct evidence against the teacher. If misalignment does not transfer, the student still inherits useful capabilities while remaining benign [1]. Either outcome is a win for safety, which is why the authors frame distillation itself, not just the techniques built on top of it, as the useful lever.
